# Mermaid architecture and benchmark specification

Version 1 · 30 September 2026 · balanced performance and energy objective · intended interface **mermaid.tide.casa**

Mermaid is the approved digital twin and engineering blueprint for a Bangel quantum semiconductor. Its first executable form combines the preserved Bangel compiler, independently admitted representation and runtime with bounded quantum and thermal models, predictive and neural software, evidence records and an interactive browser interface. Its physical target is a small silicon-spin-inspired quantum processing unit alongside classical electronics. That physical target remains a proposed design.

The project name Mermaid follows the user's decision. The preserved framework's Mermaid accounting/reconciliation role retains its own contract. The cooling work reuses **MJos Peltier–Seebeck Thermal Orchestration / MJPS Thermodynamics**, even though the recovered source uses a different name.

## System boundary

```mermaid
flowchart LR
  UI[Mermaid browser interface] --> HOST[Bounded host adapters and workers]
  HOST --> B[Bangel source]
  B --> E[Elsa compiler]
  E --> JP[Independent JP admission]
  JP --> J[Joanna execution and receipt]
  HOST --> N[Predictive, neural, recall and decision software]
  N --> B
  HOST --> Q[Mermaid quantum state model]
  HOST --> TEC[Warm-stage MJPS thermoelectric model]
  HOST --> LED[Time, provenance and energy ledger]
  J --> LED
  Q --> LED
  TEC --> LED
  LED --> UI
  Q -. proposed device mapping .-> SPIN[Silicon-spin QPU blueprint]
  SPIN -. separate qualification .-> CRYO[Cold-stage refrigeration boundary]
```

Solid arrows describe software information flow. Dashed arrows identify future physical engineering. The drawing is schematic and carries no calibrated distances, fabricated chip dimensions or demonstrated quantum advantage.

The standard inventory includes every actual computational package found in the source checkout: general Python and .NET toolchains, GIN adapters, bounded Joanna scheduling, MJ-Bangel-Neural, MJ Neural Net, MJ Memory Recall, MJos Predictive, MJ Decision, the thermoelectric reference and paper formula adapters. It also records all documented family roles, DRO/Pearl objects, the neurological umbrella and the proposed four-family cells with their implementation status. Inclusion in the inventory does not turn a documentary role into an executable processor.

## Bangel and host responsibilities

The frozen source revision is `229548f13a4856455b3d30d408cf845fb5b557fc`, from the preserved `bangel-language` checkout. Source owners' original files remain unchanged; Mermaid packages bounded local copies and separate adapters as required.

| Surface | Present capability and native seam | Mermaid responsibility |
|---|---|---|
| General Bangel 1.0 | `ElsaCompiler.compile(name, source)` or `compile_project(sources)`; `JPIR(...).verify()`; `JoannaRuntime.execute(compilation)`; `JoannaReceipt.verify()` | Preserve source/profile identity, seven runtime budgets, independent admission and original receipt. |
| GIN | `bangel.gin.evaluate(request)` generates Bangel and runs the admitted arithmetic | Preserve exact interval/clock/unit semantics; host supplies valid clock transformations and times. |
| Neural baseline | `mj_neural_net.adapter.evaluate(request, review=None, include_receipt=False)`; input `MJ-Bangel-Neural/0.1.0` | Retain bounded scoring and explicit review; distinguish lowercase research markers from runtime R states. |
| Neural overlay | `mj_neural_net.integration.overlay(...)` / `compose(...)`; request `MJ-Neural-Net/0.1.0` | Retain five-template namespace, bounded diffusion and context; no biological processor claim. |
| Memory recall | `mj_memory_recall.engine.recall(...)` | Bangel owns scoring; host owns validated catalog and SQLite persistence. |
| Predictive | `PredictiveCore.predict(data, timeout_s=15, cancel=None)`; input `PMOS-Input/1`, result `PMOS-Receipt/1` | Native host uses supervised separate workers. Browser replacement must qualify independent WebWorker execution and preserve admission/receipt behavior. |
| Decision / L1–L5 | `decide(request, checks=...)`; `IntegratedPipeline.run(request)` | Preserve independent gates, separate operating mode from decision status, and explicit millisecond cost binding. SELECTED remains advisory. |
| Paper formulas | `bangel_papers.host.evaluate(color, profile_id, params)` | Expose each formula's actual capability and evidence ceiling. Unsupported formula IDs stay refused. |
| MJPS thermoelectric | `BangelMJPSCompiler` → `JoannaMJPSRuntime`; `JP-IR/MJPS-1`; point kernels and bounded fleet optimizer | Preserve the separate Bangel 0.1 domain grammar and constant-property steady model. |
| Quantum state model | New `Mermaid-statevector-host/1` extension | Host owns complex amplitudes, gate application, noise mixture and sampling. The current general parser has no executable qubit/apply/sample grammar. Bangel can perform compatible real arithmetic and checks through explicit kernels. |
| Warm transient model | New `Mermaid-warm-RC-host/1` extension | Host owns the newly qualified RC integration; do not attribute it to the absent `mjps.network` implementation. |
| Energy and schedule models | New `Mermaid-energy-boundary-host/1` and `Mermaid-schedule-host/1` extensions | Declared external accounting and illustrative ready-job ranking; these adapters do not add native language syntax or generate forecasts. |
| Browser | Pyodide/WebWorkers or selected host stack | Transport, persistence, actual host timestamps, cancellation, UI and supported process adapters. These are host code, even when they call Bangel. |

The documentary Physics 0.1 gate listings and proposed MJ.23.R2.ACT four-family grammar remain separate references. Mermaid does not promote them to currently supported syntax. Browser code must not silently replace Bangel math with a same-named JavaScript formula and describe that replacement as Bangel execution.

## Proposed physical target and two thermal stages

The default blueprint starts with **two silicon electron-spin qubits in a double quantum dot**, with optional schematic expansion to eight. The model exposes initialization, microwave/electrical control, exchange coupling, charge/readout sensing, classical control and interconnect boundaries. Native pulse mapping, fabrication layers, layout rules, materials, gate durations, coherence times, noise spectra, readout error and calibration remain unspecified until separately supported.

Published research demonstrates silicon-spin operation above 1 K and studies integration with cryogenic CMOS. These results justify considering a silicon-spin architecture; they are external experiments and do not validate Mermaid's proposed device. The provisional cold-stage target of **1.2 K** is a scenario setting, not a measured operating point of this project. [High-fidelity spin-qubit operation above 1 K](https://www.nature.com/articles/s41586-024-07160-2), [Spin-qubit control with a milli-kelvin CMOS chip](https://www.nature.com/articles/s41586-025-09157-x).

| Stage | Model and defaults | Boundary |
|---|---|---|
| Warm electronics / thermoelectric stage | Preserved fixture: α=0.05 V/K, R=2 Ω, K=0.5 W/K; hot 320 K, cold 300 K, uncertainty ±0.2 K; current up to 2 A. Module temperature envelope 250–450 K. | Applies to the warm stage only. Coefficients and calibration flags belong to a synthetic reference fixture. |
| Cold QPU stage | Proposed silicon-spin device at 1.2 K; refrigerator cooling-capacity curve, wall-power curve, gate duration, heat leakage, dissipation and coherence parameters start **unknown**. | The warm TEC fixture cannot be extrapolated below its 250 K envelope or treated as a 1.2 K refrigerator. |
| Interconnect / packaging | Separate thermal conductance, electrical loss and stage-resolved load terms | Unknown terms retain unknown status; a trace cannot invent zero wiring heat or zero refrigeration cost. |

The thermoelectric model preserves the source sign convention: positive current pumps from the declared cold face to the hot face; positive electrical input enters the module; positive Qc enters its cold face and positive Qh leaves its hot face.

`V = α(Th−Tc)+IR`

`Pin = VI`

`Qc = αTcI−I²R/2−K(Th−Tc)`

`Qh = αThI+I²R/2−K(Th−Tc)`

`Qh−Qc−Pin = 0`

For the reference point at 2 A, the nominal fixture expects 16 W cooling, 26 W hot-side heat rejection, 10 W pump input and COP 1.6. Four temperature corners give 15.78 W guaranteed cooling and 10.04 W worst-case pump input. These are mathematical fixture expectations. The source hybrid example uses two distinct modules; it accounts for harvested power at converter efficiency 0.95, pump input at driver efficiency 0.9 and auxiliary power 0.25 W. Heat recovery does not erase those costs.

Calibration and missing-input gates are part of admission. The source's full `JoannaMJPSRuntime` checks calibration. Raw `coupled_kernel` or `optimize_fleet` calls do not themselves provide that entire receipt path; direct Mermaid adapters must preserve its required gate behavior and identify themselves accurately.

The transient extension uses a warm node with heat capacity C J/K, thermal resistance R K/W, supplied heat-load trace and declared passive boundary: `C*dT/dt = Pload−(T−Tamb)/R−Qcool`. It records each step's stored-energy change, incoming heat, passive loss and cooling extraction. Time-step bounds, convergence, temperature-envelope violations and energy residuals are qualification requirements. Its optional policies are no active TEC, fixed/reference TEC, robust demand-based TEC, and distinct harvester plus robust TEC. No policy may quietly count the same physical module simultaneously as independent generator and pump.

## Versioned state and receipt contracts

The JSON contracts live in `work/mermaid-contracts/`: `common.schema.json`, `runtime.schema.json`, `quantum.schema.json`, `cooling.schema.json`, `energy.schema.json`, `schedule.schema.json`, `defaults.json`, `processor-manifest.json`, `benchmark-suite.json`, example requests and matched cooling presets. They are integration artifacts owned by this specification; deployment may expose reviewed copies later.

All wrappers use a declared schema version, request identity, bounded workload and a named source profile. Native payloads are validated by the actual native adapter; wrapper validity does not establish native validity.

- A quantity carries value, unit, evidence class, origin, uncertainty, source identity and calibration identity. Unknown is `null`; it is distinct from zero. Classes are **modelled**, **predicted**, **observed** and **measured**. A received synthetic value can have an observed receipt time while its value remains modelled.
- Provenance includes exact source revision, input/program hashes when known, original source IDs, parent receipts and evidence status. A content hash binds content; it is not authentication or physical measurement.
- Event times preserve expected, acquired, observed, recorded, delivered, reconciled and completed milestones. Each time uses a clock identity, unit, interval endpoints, temporal state and admitted transform identity. Different clocks without an admitted mapping stay incomparable.
- Every receipt has status, named reasons, outputs, provenance, event times and budget usage. Software effects remain NONE and execution permission false. Native receipts are retained; receipt verification cannot be replaced by a status badge.
- Error/HOLD/cancellation preserves its actual disposition. A successful process exit or correct scalar predicate does not override a refused request. Partial model diagnostics are explicitly marked and never presented as admitted physical results.

Three clocks remain separate: browser monotonic elapsed time, the scenario's modelled schedule clock, and any future instrument/device clock. Simulator operations, neural nodes and receipt lifecycle stages do not define physical clock cycles, qubits, transistor counts or joules.

## Bounded backend seams

The host gateway calls `evaluate_quantum(request)`, `evaluate_cooling(request)`, `evaluate_energy(request)` and `evaluate_schedule(request)` and wraps or retains their versioned receipts. General runtime calls route through the preserved APIs. The host implementation language is separate from the data contract.

Quantum request v1: 1–8 qubits, initial all-zero state, little-endian basis with q0 as the least significant bit, H/X/Z/CX/CZ/PHASE gates and explicit whole-register RESET_ALL, final independent bit-flip noise or none, explicit measurement draws in [0,1), and a bounded gate/shot/cell-update budget. CX/CZ require distinct two-qubit targets; PHASE alone takes a finite angle. Result probabilities include the declared noise; optional ideal amplitudes are labelled ideal. No unsupported gate or budget violation yields success-looking counts.

RESET_ALL targets must name every qubit once. It implements the valid full-register replacement channel to the all-zero state, including after entanglement; partial reset and mid-circuit selective measurement are unsupported. Reset operations consume a gate/event and bounded state-cell updates. The reset environment and physical reset energy remain outside the current model.

Optional `timingModel` is null/unknown by default, or the closed `mermaid.quantum.timing-model/1` profile in ns. Its gate-duration map covers H/X/Z/CX/CZ/PHASE; `readoutNs`, `resetAllNs` and `classicalHandoffNs` are explicit decimal strings or null. Output `modelledTimeline` retains gate/reset rows, a compressed READOUT_BATCH with explicit repetitions, and CLASSICAL_HANDOFF. A missing duration makes its end and later absolute starts unknown; known-duration subtotal does not close the timeline. It models one source circuit plus independent samples of the final distribution; hardware per-shot re-preparation, real pulse timing and device speed remain unmodeled. Host monotonic receipt timestamps are a separate observed clock.


Cooling request v1: warm-stage module declarations, expected/anticipated/delivered temperatures, uncertainty and calibration, mission demand, and point/optimize operation. Native field names and meanings are preserved. A bounded transient operation uses its explicitly separate warm-RC model and load trace. Module IDs and observations must match; nonfinite numbers, unknown delivered temperatures, unsupported mode, uncertain/reversed orientation, envelope violations and unverified calibration produce a named refusal or HOLD.

Schedule request v1: at most 64 chronologically ordered model jobs, two homogeneous logical workers, explicit task duration/active-power assumptions, two worker idle powers, explicit per-dispatch prediction/neural overhead and supplied advisory records. Its default S01 trace has 64 jobs in 16 bursts of four; the task cycle is general Bangel, quantum, predictive and neural overlay. The illustrative durations are 5, 8, 12 and 4 ms and active powers 2, 2.5, 3 and 1.5 W; each worker's idle power is 0.3 W, prediction attempt cost is 0.5 ms, neural attempt cost is 0.25 ms and their active power is 1 W. These inputs are scenario assumptions, not benchmark observations.

The four model policies are FIFO fixed, prediction only, neural only and combined. Performance uses ascending supplied estimated-duration rank; energy uses ascending estimated-duration × supplied active-power rank; balanced uses both ranks. Neural uses descending supplied priority rank from an explicitly identified caller projection. Combined averages the applicable ordinal ranks. Equal scores retain trace order. Each dispatch falls back to FIFO when any ready job lacks required currently available advice; future advice is excluded. The default supplies no advice and discloses that fallback. Configured attempt overhead remains charged on fallback. Native receipts remain the caller's evidence; this model neither synthesizes native predictions nor validates a neural-to-scheduler projection.

Schedule outputs include every model timeline, completion-set parity, nearest-rank p50/p95, makespan, throughput, active/idle/overhead joules and model comparisons. Task durations and energies remain identical across policies; order and attempt overhead may differ. Model completion does not execute a Bangel program or establish result correctness. Actual host execution and the approved host-p95 target require separate measured paired runs. Its host energy subtotal excludes warm electronics, cold QPU and refrigeration and therefore leaves the whole-system energy target unassessed.

Default limits: two host workers, 10 s wall timeout, 8,192 wrapper events, 1 MiB wrapper output; quantum default 64 gates, 1,024 saved midpoint draws, maximum 4,096 shots and 1,000,000 state-cell updates. Absolute request limits are eight qubits and 512 gates. Native Joanna default limits are 4,096 instructions, 8,192 events, depth 64, 16,384 evaluations, 1,000,000 allocations, 1 MiB output and 4 MiB receipt. Compatibility profiles that lack custom-limit support stay explicitly refused. Workers return cancellation evidence; the browser host enforces hard interruption when cooperative cancellation is insufficient.

The numerical schema fields use finite decimal strings where the preserved runtime expects them. JSON schemas define shape; semantic checks additionally enforce units, bounds, target count, profile matching, array length, finite numbers, monotonic time, calibration and ledger conservation.

## Power, thermal and energy accounting

The ledger records electrical input/output, heat input/output, stored-energy change, elapsed duration and residual at each named boundary. Watts are rates and joules are accumulated energy. For a sampled signal use a declared integration rule; for a scenario use a declared piecewise-constant model. Neither mode converts allocation counts or benchmark scores into joules without a supplied energy model. A known electrical subtotal alone does not close the ledger: missing heat or stored-energy terms retain unknown conservation status, null full total and unassessed whole-system target.

Internal transfers are recorded for conservation but are not repeatedly summed into total external consumption. The total comes from unique external boundaries. Compiler/runtime, both predictive workers, independent review, neural overlays, persistence/transport, UI, idle time, TEC driver/converter loss, auxiliaries and refrigerator wall power are included when active in that boundary. Energy, electrical charge, bytes, model scores and money remain separate accounts.

The initial energy subtotal names its active envelope **Mermaid-active-software-and-warm-stage/1**. Quantum state simulation is host computation in that envelope. The proposed physical QPU and cryogenic plant are not operating hardware. The **whole-system target** includes its cold-stage and refrigeration model boundaries: unknown quantities leave that total **partial / incomparable**, with known subtotals available and the whole-system improvement unassessed. A warm-stage subtotal cannot replace the approved whole-system target. UI assumptions are labelled and versioned. Missing energy terms never become zero.

Distinct harvester and pump reservoirs remain separately named in cooling outputs. Hot-side pump rejection is not reduced by unrelated generator heat extraction; the generator's hot-source extraction and cold-sink rejection are separate fields. Transient summary power comes from the final complete trajectory step, including the distinct generator, while accumulated joules come from the entire trace.

## Frozen benchmark suite and balanced objectives

The approved performance target is **15% lower actual host end-to-end p95 latency** and **20% lower modelled whole-system external joules per completed job**, against a fixed policy on the same workload. These are development objectives, not results. Actual browser/host measurements include all added workers and adapters; an unmet target is reported as unmet. Modelled schedule latency and the known warm-stage energy subtotal are useful additional metrics, and cannot substitute for either approved target. Whole-system energy remains unassessed while cold-stage or refrigeration inputs are unknown.

| Family | Fixed case | Expected evidence |
|---|---|---|
| Bell state | Two qubits, H(q0), CX(q0,q1), 1,024 midpoint draws | Probabilities [0.5,0,0,0.5]; 512 counts each of 00/11; none of 01/10. |
| Independent bit flips | Same Bell circuit, p=0.1 before measurement | Probabilities [0.41,0.09,0.09,0.41]; mismatch probability `2p(1−p)=0.18`. |
| Small search | Two-qubit Grover recipe, marked state 11 | Unit probability on 11 within numerical tolerance. This is a small simulated algorithm test. |
| Quantum refusals | Invalid draw, target, gate, noise, aliasing or budget | Specific refusal; no fabricated outcome. |
| Forecast | Preserved duration fixture plus chronology/context/unit failures | Exact native contract parity; chronological later outcomes for accuracy and coverage evaluation. |
| Neural / memory | Partial cue, five-template overlay, entity/context conflict, unknown and near tie | Preserved native scores/disposition; wrong context excluded and ambiguity retained. |
| Decision | Admitted candidate, missing review, independent denial, expiry and repeated no-progress loop | Advisory status and gate precedence preserved; permission remains false. |
| Steady thermal | Point and source hybrid bus fixture | Pump/cooling/rejection/COP and four-corner bounds; full electrical cost and conserved ledger. |
| Transient thermal | Constant load analytical RC case, changing load, time-step refinement, envelope crossing | Convergence against analytical solution, integrated energy residual, bounded violation status. |
| Timing | Point/interval/unknown clocks, admitted transform and incomparable clocks | Preserve native GIN semantics; no inferred physical calibration. |
| Mixed scheduling | 64 jobs; bursts of four every 10 modelled ms; fixed seed 2309; two workers | Compare fixed, predictive only, neural only and combined strategies. Workload completion/correctness and all added overhead are matched. |
| Cooling ablation | Same mixed trace, ambient, heat load, limits and accounting boundary | Compare no-active-TEC, reference-fixed-TEC, robust TEC and separate harvester plus robust TEC; thermal feasibility is a guard. |

Five warmup runs precede 30 recorded runs in each of three independent sessions. Alternate baseline/candidate order; retain every valid run and all failed/held/skipped results. Save draw arrays and workload seed. Report cold initialization separately from warm execution. p95 uses nearest rank `ceil(0.95*n)` on a declared sample set; do not pool different workload units into one undocumented score.

Report p50/p95, throughput, completed/requested denominator, coverage/refusal reasons, maximum temperature, settling behavior where qualified, residuals, known and unknown energy terms, modelled J/completed job and raw artifacts. Forecast evaluation reports MAE, signed bias and answered/requested coverage on later disjoint cases. Interval coverage is not assumed calibrated. The first synthetic held-out challenge freezes four training outcomes, eight later disjoint outcomes, one horizon and quality thresholds before execution. Recovered MJos/Teal outputs and Bangel reconciliation/summary scores are compared with independently computed host arithmetic, last-observation and moving-average baselines. Chronology/context/unit/evidence gates and downstream review HOLD remain in force. Quality verdicts describe that supplied synthetic series, not device prediction accuracy.

Repeated modeled ablations use three fresh interpreter sessions, five warmups and30 measured runs per scenario/objective, with all four policy outputs retained. No-advice FIFO fallbacks and explicitly synthetic prior advice are separate scenarios. Cooling policies share the same installed modules, ambient, heat-load trace, initial state and temperature limits; step refinement checks numerical convergence. Active/idle/advisory and warm external energy are reported with separate ledgers and fixed comparison windows. Neither repeated model evaluation host timing nor a known subtotal establishes the approved native mixed p95 or whole-system energy target. Scheduling, recall, independent gates and cooling each receive separate ablation results.

An objective is met only when correctness and completion are matched, no new thermal violation occurs, no unavailable component or unknown energy term is hidden, and the specified paired comparison supports it. Otherwise report **not met**, **incomparable** or **unverified**. Improving a simulated small circuit does not establish physical quantum speedup.

## Verification and later physical qualification

Schema/reference/fixture checks establish that the integration specification is coherent. Fresh unit and integration tests must qualify the numerical backends, native receipt verification, profile refusals, browser worker cancellation, persistence and UI behavior. Native-source parity and browser adapters require separate results. No completion badge inherits an old test count.

The physical blueprint needs device identity, fabrication/process rules, wiring and magnetic-field plan, calibrated temperatures, refrigerator heat-load/wall-power curves, native gate/pulse compilation, actual job submission and measurement records, error/leakage analysis, and matched physical benchmarks before it can be described as validated hardware. Imported instrument data is a distinct future adapter. This phase preserves the path to that work without manufacturing its evidence.

Source anchors: semiconductor planned-only product BP-PROD-037, encyclopedia master p2423; thermal formula TR-PINK-F-152 p4046; wake reserve F-233 p4130; MJPS Thermodynamics Pink p2005 / Teal p2974 / Indigo p7984; book chapters 17–20,25–26,31,36–40,46. Exact JSON IDs and source statuses are retained in this chat's research map.
